Papers with forced alignment

6 papers
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement (2025.acl-long)

Copied to clipboard

Challenge: GigaSpeech 2 is a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages.
Approach: They propose a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages and an automated pipeline for data crawling, transcription, and label refinement.
Outcome: The proposed corpus reduces the word error rate for Thai, Indonesian, and Vietnamese on a realistic YouTube test set by 25% to 40% compared to Whisper large-v3.
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented.
Approach: They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods.
Outcome: The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents.
Extracting Linguistic Knowledge from Speech: A Study of Stop Realization in 5 Romance Languages (2022.lrec-1)

Copied to clipboard

Challenge: voicing alternation phenomena of stops are a common problem in connected speech . phoneticians and phonologists are interested in analyzing phonetic variation .
Approach: They use forced alignment with pronunciation variants and machine learning techniques to examine voicing alternations of stops in Romance languages.
Outcome: The proposed method enables linguists to use large corpora and speech recognition systems . the results show that voicing alternations occur in all Romance languages .
Structured Tree Alignment for Evaluation of (Speech) Constituency Parsing (2024.acl-long)

Copied to clipboard

Challenge: Recent work has proposed a new task of textless speech constituency parsing that uses textless parsers to parse spoken word boundaries over automatically recognized spoken word borders.
Approach: They propose a metric that compares a constituency parse tree over spoken word boundaries with a ground-truth parser tree over written words.
Outcome: The proposed metric shows higher tolerance to syntactically plausible parses than PARSEVAL.
A Romanization System and WebMAUS Aligner for Arabic Varieties (2022.lrec-1)

Copied to clipboard

Challenge: The WebMAUS 1 is a suite of webservices that is free for academic users that processes 42 languages and language varieties.
Approach: They propose to develop an Arabic variety-independent romanization system that aims to homogenize and simplify the romanization of the Arabic script.
Outcome: The proposed system is based on the existing Arabic variety-independent WebMAUS services.
English-based acoustic models perform well in the forced alignment of two English-based Pacific Creoles (2025.acl-long)

Copied to clipboard

Challenge: Currently, European languages dominate phonetic research . forced alignment can accelerate the study of sociophonetic variation in minority languages .
Approach: They propose to use English and custom-made acoustic models to study the alignment of vowels in two Pacific Creoles, Tok Pisin and Bislama.
Outcome: The proposed models perform acceptablely well in English and humans in vowel environments described as ‘Highly Reliable’.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations